A statistical framework for evaluating the repeatability and reproducibility of large language models
This paper presents a regulatory-informed statistical framework that quantifies the semantic and internal repeatability and reproducibility of large language models, demonstrating that these metrics vary significantly based on prompting strategies and model configurations, are often independent of diagnostic accuracy, and are essential for systematically evaluating LLM reliability in biomedical applications.